Method & standards

The rules this build enforces, the standard a Core sense line is held to, and how each of them is checked.

Scope

This document states the rules the build enforces and how they are checked. It is the companion to the reader's key, which explains what the pages mean. Every figure here is read from the build it ships with.

1,604
roots built
25,845
entries
60,139
senses
175,637
Arabic spans placed

1 · The rules that do not bend

Arabic never passes through a language model
Every Arabic span is lifted out before the text is sent, replaced by a numbered token, and substituted back at render time. The tokens are counted on the way in and on the way out: the build fails if one is lost, and fails if one is used twice. This is why no Arabic in the build can be paraphrased, normalised or invented — it never reaches a model to be altered.
Attested only
No verb form, no word form and no meaning is generated to fill a gap. Where the source is silent the page is blank. A sense whose source text carries no words at all renders as nothing, and the build hard-fails if such a sense comes back carrying content.
Entries and senses are Lane's divisions
Read off his own text. Nothing may add, drop, merge, split or reorder one; the gate asserts the entry count and the full sense sequence against the source, and refuses the build on any mismatch.
Refine, never regenerate
Text is repaired in place. Rebuilding a passage from scratch destroys work that was already correct, and it has: an early attempt regenerated core lines and lost 101 of 131.
Never a paraphrase alone
Every entry keeps a toggle to the verbatim source text.
Order is evidence
Glosses and notes stay in Lane's source order. He writes a note directly after the phrase it comments on, so for a note carrying no Arabic its position is the only evidence of what it belongs to. Both arrays are checked, and a reordering is a build failure.

2 · Identity: what joins, and what is shown

Each sense carries two numbers and they are not interchangeable.

PurposeShape
idThe join key. Immutable. Errata, provenance and every internal reference match on it.flat — 4.1, 4.2
displayWhat the reader sees. Carries Lane's own two-level numbering.hierarchical — 1.1, then 1.1.1

Anything reader-facing shows display; anything that joins uses id. Renumbering id would silently break every join, so it is frozen. This build has 33,159 top-level senses and 26,989 subordinate.

3 · The Core sense standard

The boundaries below are final. They are carried in the generation prompt, not applied afterwards as a filter.

Up to three distinct meanings — a recommendation, not a limit
Rarely more. Different shades of one meaning count as one: a mark, a sign, a token has used none of the allowance. This raises the ceiling in practice rather than lowering it.
Priority when meanings compete
(a) closeness to the entry's core meaning; (b) position in Lane's order, as the tie-break only. That ordering is load-bearing: شهد's "he was present at it" is Lane's last sense, and position-first would have dropped the very meaning that matters most.
Length is not a criterion, anywhere
Deliberately. Volume misleads in both directions — one entry lost a meaning about captivity to a tree that filled 46% of the article, and another lost a sense worth 21% of its entry. A meaning that takes one line to state is not thereby less central.
Omission is the graver error
Tighten the wording before dropping a meaning.
The core line's form follows the entry's
A nominal entry gets a nominal line; a verb entry gets "he struck him", not "to strike". A verbal line on a nominal entry is reported for review.

Retired, and not coming back: the build no longer marks individual words in a core line as Lane's own, and there is no second pass that rewrites core lines. An audit of all traceable meanings measured a fabrication rate of zero, so the mark guarded nothing while itself introducing errors. 333 marked words remain in data from earlier builds and are no longer shown.

4 · Vocabulary policy

A settled replacement never comes back. The list is applied mechanically and then enforced — a root arriving with one of Lane's retired words fails the build rather than being quietly fixed, because that would hide a generation problem.

Words are replaced only where a single substitution is always right. Terms of art are left alone: predicament stays where Lane means Aristotle's category, and a word whose plain equivalent is a phrase rather than a word is not substituted, because that changes his sentence rather than his vocabulary.

5 · The Qur'ānic join

What an occurrence claims
Only that the word occurs in that verse. Never that it illustrates a particular sense — that claim produced a wrong attribution early on and was removed. Occurrences attach to the entry, never to a sense or a gloss.
How an entry is matched
A verb entry matches on the corpus's own form tag. A noun entry matches Lane's vowelling against the corpus lemma, same root, same part of speech. Two possible matches means show nothing.
Verse numbering
Lane cites Flügel; modern print uses the Kufan numbering. The mapping is built and self-checking, and a citation that cannot be mapped with certainty is left as Lane wrote it.
An āya is quoted only when
the citation is direct — never a tafsir reference — and the verse actually attests the root. When in doubt, nothing is quoted.

The word map, in the Qur'ān Explorer

Each word of the Qur'ān is linked to the root Lane files it under. A word map refuses rather than shifts: if the reader's word count and the corpus's disagree by one, every word after that point resolves to the wrong root and nothing looks broken, so a verse that cannot be aligned gets no map at all. The map emits one slot per displayed token, never per corpus word, and keeps a separate array per orthography.

It is verified at 99.8% — and, more usefully, the same measure collapses to about 3.5% when the alignment is deliberately shifted by one word. That collapse is what makes the number mean anything.

Particles: a link that claims less

A third of the Qur'ān's words carry no root at all — every preposition, pronoun, relative, negator and conditional, 27,462 words in all, and only 175 different lemmas among them. Lane writes articles for them; he simply files them under a bare letter rather than a root, so مَا sits under م and لَا under ل. A root-based join could never see them.

These links claim something weaker than a root link, and the difference matters. A root link says this is the root of this word. A particle link says only Lane discusses this word in this article — never that the letter it is filed under is its root, because م is not the root of مَا. The Qur'ān Explorer says so on hover, and the published map keeps the two in separate tables so nothing can read one as the other.

Which article a word belongs to is stated one word at a time, by hand, never matched. Matching on the bare letters cannot work: من is three different articles in Lane — مِنْ the preposition, مَنْ the relative, and مَنٌّ a favour — and a matcher would send thousands of words to whichever it happened to find first. So there is an enumerated table of all 175 lemmas, each row read by eye, each carrying its evidence. A lemma with no row links to nothing; it is never guessed at. That is why ٱلَّذِى, the single commonest of them at 1,464 words, is deliberately left unlinked: Lane has no article on it, the nearest is ذا, and two possible answers means showing none.

6 · Traditions

Lane cites roughly 1,600 hadith across the corpus. He gives a collection, book or number for none of them — 0%, against 83% of his Qur'ānic citations, which carry sura and verse and are joined up for the reader.

So the red on a matn claims exactly one thing: that Lane attributes this Arabic to a tradition. It is not a citation and is not joined to any hadith corpus. Matching a matn fragment by text alone would be a weaker join than the project's rules allow, and would claim more than the source supports.

Only multi-word Arabic is marked, and only where a note attributing it to a tradition is attached to it: a matn is a sentence, while a single Arabic word beside it is one of Lane's glossed forms. 1,900 passages across 781 roots.

7 · Errata

14 corrections to the electronic source, and the policy behind them:

Enumerated, never a pattern
Each names its root, entry, sense and the exact text expected, and carries its evidence. Pattern rules destroy real words — one root genuinely contains a form a plausible rule would have "corrected".
Exactly one match, asserted
A correction whose target no longer matches exactly once fails the build rather than silently doing nothing.
Shown, never silent
The reader sees what the source said beside what it was corrected to, and the verbatim text is left uncorrected so the original is always visible.

8 · How the build is checked

One command runs the whole chain and returns a single verdict. The checks that must read zero: Arabic lost, Arabic duplicated, unknown token, empty sense, marker leak, order mismatch. Arabic placed plus Arabic in the morphology bars must reconcile to the total the source contains.

Two frozen reference points
A ten-root set signed off by hand, and a fifteen-root standard built on it. Their entry counts, sense counts and rates are asserted on every run. Neither may move; if either does, something has changed that was agreed not to.
A guard is not trusted until it has been seen to fail
This is the rule that matters most. The build deliberately breaks its own checks on every run — planting literal Arabic in the prose, a Roman numeral, a retired vocabulary word, an out-of-range reference — and requires each one to be caught. A check that cannot fail is worse than none, because it is believed. This project has shipped two such checks and found them only by trying to break them.
A guard must not share its logic with the thing it guards
The same test is deliberately written out more than once, in different files, even where that repeats a few lines. A shared helper lets one bug silence both the fix and the check together — which has happened here, and stayed green.
Metrics are not sufficient
Every serious defect in this project's history was invisible to the counters. So the build ends by naming the places a fault is most likely to hide, for a person to read. That list is capped and ranked, because an exhaustive dump is the same as no list.

9 · The published data is the product

The HTML explorer is one consumer of the data, not the deliverable. The published data carries Arabic substituted in, markers as structured fields, a schema version and a build manifest. Anything emitted only in rendered form, or in a shape that leaks the pipeline's internals, has failed this rule — and the build checks that the published package contains none of them.